Papers with natural language processing tools
Automatic Discovery of Heterogeneous Machine Learning Pipelines: An Application to Natural Language Processing (2020.coling-main)
Copied to clipboard
| Challenge: | Existing AutoML systems use heterogeneous techniques to build pipelines that combine techniques and algorithms from different frameworks. |
| Approach: | They propose a system for automatic machine learning that uses heterogeneous techniques. |
| Outcome: | The proposed system is evaluated in diverse machine learning problems and compared with other alternatives. |
COIN – an Inexpensive and Strong Baseline for Predicting Out of Vocabulary Word Embeddings (2022.coling-1)
Copied to clipboard
| Challenge: | Word embedding models only include terms that occur a sufficient number of times in training corpora. |
| Approach: | They propose a method for predicting word embeddings for out of vocabulary terms using word2vec. |
| Outcome: | The proposed method surpasses several methods on benchmark tasks and is inexpensive to compute. |
Creating Terminological Resources in the Digital Age for Less-resourced Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | Multilingual terminological resources are limited in less resourced languages, limiting knowledge spread in less-resourced languages . linguists and terminologists must use natural language processing tools to maximize resources . less-represented languages suffer from a lack of available linguistic resources - a survey shows . |
| Approach: | They propose a method to maximize the open access catalan terminology available . authors propose linguists and terminologists supervise the project and translate it into catalane . |
| Outcome: | The proposed method maximizes the catalan terminology currently available in open access . the results are supervised by linguists and terminologists experts before being publicly available to the public. |
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)
Copied to clipboard
| Challenge: | Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools. |
| Approach: | They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities. |
| Outcome: | The proposed method matches dialect areas at different granularities against an existing dialect map. |
BILinMID: A Spanish-English Corpus of the US Midwest (2022.lrec-1)
Copied to clipboard
| Challenge: | The Hispanic population in the United States is up to 15-20% of the nation's total population . due to its proximity to the US-Mexico border, Hispanicas have more presence in the Southwest of the country . |
| Approach: | They propose to create a text corpus of the Spanish and English spoken in the US Midwest by different types of bilinguals. |
| Outcome: | The proposed corpus contains short stories narrated in Spanish and in English by 72 speakers representing different types of bilinguals: early simultaneous bilinguals, early sequential bilinguals and late second language learners. |
CoSimLex: A Resource for Evaluating Graded Word Similarity in Context (2020.lrec-1)
Copied to clipboard
Carlos Santos Armendariz, Matthew Purver, Matej Ulčar, Senja Pollak, Nikola Ljubešić, Mark Granroth-Wilding
| Challenge: | Existing methods to evaluate word embeddings ignore context and treat words in isolation. |
| Approach: | They propose to build a new word embeddings-based dataset that provides context-dependent similarity measures. |
| Outcome: | The proposed dataset provides context-dependent similarity measures and covers a well-resourced language (English) but a number of less-resource languages. |